Machine Learning
Machine learning in health is less about model architecture than about data, evaluation and deployment. The modelling is usually the easiest part of the problem.
Task types
| Task | Health example |
|---|---|
| Binary classification | Will this patient be readmitted within 30 days? |
| Multi-class classification | Which of these conditions does the presentation match? |
| Regression | Predicted length of stay |
| Time-to-event (survival) | Time to treatment failure |
| Time series forecasting | Expected outpatient volume; commodity demand |
| Anomaly detection | Unusual reporting patterns in surveillance data |
| Clustering | Patient segmentation for programme design |
| NLP | Extracting structured facts from clinical notes |
Where the difficulty actually is
Label quality. Diagnosis codes are recorded for billing and reporting, not for model training. A "diabetes" label may mean confirmed diagnosis, suspected diagnosis or a rule-out that was never removed.
Missingness is informative. A test that was not ordered tells you something about the clinician's judgement. Naive imputation discards that signal — and can inject it as leakage.
Label leakage. Features that encode the outcome (a treatment only given after diagnosis, a discharge-time field) produce excellent validation scores and useless deployed models.
Class imbalance. Serious outcomes are rare. Accuracy is meaningless; use precision-recall, and choose thresholds against clinical cost.
Distribution shift. Case mix, protocols, coding practice and populations change. A model trained on last year's data quietly degrades.
Representativeness. Data from urban tertiary hospitals does not describe rural primary care. Models trained on one and deployed on the other fail on the people who are least well served already — see AI ethics.
Evaluation
- Split by time and by site, not randomly. Random splits leak future and site-specific information.
- Report calibration, not only discrimination. A well-ranked but miscalibrated risk score misleads decisions.
- Use decision-analytic metrics — net benefit, cost-weighted error — because false positives and false negatives are not equally harmful.
- Report subgroup performance by sex, age, geography and other relevant strata.
- Compare against the real baseline: current clinical practice or the existing rule, not against chance.
Deployment
A model in a notebook has changed nothing. Deployment requires:
- An integration point — where the prediction reaches the decision-maker
- Latency and availability guarantees appropriate to the workflow
- Monitoring for input drift, output drift and performance decay
- A rollback path
- Governance: who is accountable for the prediction, and what happens when it is wrong
See clinical AI for the regulatory and safety dimension, imaging AI for the imaging-specific path, and health data for what you are allowed to train on.